Back

European Radiology

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match European Radiology's content profile, based on 15 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Assessment of the accuracy of lung lesions diagnosis in adolescents with osteosarcoma using artificial intelligence

Uskova, N. G.; Gombolevskiy, V. A.; Chernina, V. Y.; Burenchev, D. V.; Akhaladze, D. G.; Panina, E. V.; Karachunskiy, A. I.; Tereschenko, G. V.; Goncharov, M. Y.; Soboleva, E. A.; Konopleva, E. I.; Bydanov, O. I.; Plekhov, S. Y.; Grachev, N. S.

2026-06-10 radiology and imaging 10.64898/2026.06.08.26354011 medRxiv
Top 0.1%
28.2%
Show abstract

Background. Lung metastases in osteosarcoma (OS) are the main cause of the death. The accuracy of the diagnosis of nodules by computed tomography (CT) of the lungs is critically important for determining the disseminated stage of the disease and planning surgical treatment. The use of artificial intelligence (AI) in the search for lung nodules increases the accuracy of diagnosis and reduces the chance of missing metastases. Objective: to evaluate the accuracy of lung nodules diagnosis in adolescents with OS using AI. Methods. A retrospective assessment of CT scans of adolescents with OS was performed. A pathological nodule with an average size of [≥]4 mm was considered a target finding. The diagnostic accuracy of an AI algorithm previously trained on an adult dataset was evaluated, and the number of false positives (FP) and false negatives (FN) was determined. Sensitivity, specificity, accuracy, area under the ROC curve (AUC), positive predictive value, negative predictive value, and F1-measure were calculated. Based on the obtained results, the effectiveness of the algorithm was assessed. Results. 248 CT scans of adolescents with OS were evaluated. The following results were obtained: in 5 cases, the AI algorithm showed a FP result (2.02%), in 34 cases, it showed a FN result (13.71%), and in 209 cases, a correct result (both true positive and true negative) (84.27%). The diagnostic accuracy of the algorithm was 0.843 (95% CI 0.794-0.887). The application of the AI algorithm in the practice of an X-ray doctor in a specific clinical task would allow to increase the sensitivity from 0.805 to 0.891, while ensuring an absolute decrease in the number of FN results by 8.59% and a relative decrease by 44%. Conclusion. The obtained results confirm the practical value of the application of the AI algorithm and justify the implementation of AI-assisted systems in the diagnostic protocols for lung metastases in adolescents with OS.

2
Automated AI-Based Ventricular Subcompartment Segmentation and Volumetry in Idiopathic Normal Pressure Hydrocephalus

Mutke, M. A.; Griot, S. A.; Wasserthal, J.; Indrakanti, A. K.; Vishwanathan, N.; Mahmutoglu, M. A.; D'Antonoli, T. A.; Bach, M.; Psychogios, M. N.; Lieb, J. M.

2026-06-15 radiology and imaging 10.64898/2026.06.14.26355627 medRxiv
Top 0.1%
23.2%
Show abstract

Purpose In idiopathic normal pressure hydrocephalus (iNPH), longitudinal monitoring of ventricular size is important for diagnosis and treatment follow-up. This study aimed to validate a fully automated AI model for CT ventricular volumetry with subcompartments and to compare AI-derived volume changes with routine radiology assessments. Methods This retrospective, single-center study included 88 patients with iNPH and 456 non-contrast-enhanced head CT examinations. The model was trained on 38 manually labeled CT scans with 12 ventricular subcompartments. Outcomes included segmentation accuracy, correspondence between AI-derived longitudinal ventricular volume changes and radiology report categories (decreased, unchanged, increased), radiologist detection thresholds for ventricular change, and paired pre- and postoperative volume changes in 22 patients with ventriculoperitoneal shunt. Results Mean segmentation accuracy was high (Dice, 0.83). 91% of 100 segmentations were rated as excellent by an expert neuroradiologist. AI-derived ventricular volume changes corresponded well to radiology report categories (median total ventricular volume changes of -17% in cases reported as decreased, 0% in unchanged cases, and +22% in increased cases; all p < 0.001). Radiologists reported ventricular volume change in 50% of cases at an AI-measured relative volume change of +/-6%, and in 90% of cases at +21% for enlargement and -18% for decrease. After shunt placement, ventricular volume decreased by -8% (median), with the largest relative reductions observed in the right temporal and occipital horns. Conclusions Automated AI-based ventricular segmentation on CT enables accurate and reproducible assessment of ventricular volume changes in iNPH and complements routine radiological evaluation for longitudinal and postoperative monitoring.

3
Prediction of Subsolid Pulmonary Nodule Evolution from Baseline CT Using Temporal Imaging Models

Bondarenko, M.; Qi, K.; Nowroozi, A.; Kim, J.; Kunzang, B.; Lee, A.; Liu, J.; Tran, N.; Weng, S.; Vella, M.; Chaudhari, G.; Schnizler, T.; Innanje, A.; Chen, T.; Sohn, J. H.

2026-08-13 radiology and imaging 10.64898/2026.08.12.26360292 medRxiv
Top 0.1%
22.7%
Show abstract

Background: Prediction of subsolid pulmonary nodule (SSN) progression from baseline CT may improve risk stratification and surveillance planning, but prior approaches have largely relied on fixed follow-up intervals. Methods: This retrospective single-center study evaluated interval-aware temporal imaging models for predicting future SSN growth and morphology across heterogeneous surveillance durations. A total of 24,946 longitudinal scan pairings derived from 2,543 clinician-reviewed SSNs in 426 patients were analyzed. A discriminative deep learning model predicted interval growth from baseline CT, segmentation masks, and interscan interval information, while a temporally conditioned generative model predicted future lesion morphology. Results: The discriminative model achieved an area under the receiver operating characteristic curve of 0.772 (95% confidence interval: 0.704-0.818), with sensitivity of 80.2% and specificity of 58.7% on the test cohort. The generative model predicted future lesion morphology with a Dice similarity coefficient of 0.706 +/-0.186. Prediction performance decreased with increasing follow-up duration, although both models generalized across intervals ranging from months to years. Conclusion: Interval-aware temporal imaging models enable the prediction of future SSN growth and morphology from baseline CT while accounting for variable surveillance intervals. These findings suggest a framework for time-aware, personalized risk assessment that may support individualized surveillance strategies and future AI-assisted management of pulmonary adenocarcinoma spectrum lesions.

4
Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Al-Hebshi, S.; Khalifa, H.; Pham, T. D.

2026-08-12 dentistry and oral medicine 10.64898/2026.08.11.26360189 medRxiv
Top 0.1%
19.2%
Show abstract

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings text can also be generated directly by a large language model from the image itself--raising the question of how much diagnostic value such AI-derived text carries, and whether that value depends on independent verification. Multimodal artificial intelligence (AI) benchmarks risk overstating performance if the provenance of each input--image, raw AI-generated text, or radiologist-verified text--is not clearly separated and reported. Methods: We used 300 mid-sagittal CBCT slices from the MMDental dataset. ChatGPT generated findings text and a provisional normal/abnormal label for every slice (majority vote, three independent readings from the image alone); primary classification performance was assessed on this full, unfiltered set (n=300). A radiologist then independently reviewed each case's image together with ChatGPT's description, producing their own diagnosis; three cases were excluded as insufficient, yielding 297 confirmed cases. On this subset, every model was retrained and re-evaluated under identical 10-fold cross-validation on both the provisional ChatGPT-only labels ("pre") and the radiologist-confirmed labels ("post"), isolating the effect of label provenance from image or architecture. Eight vision architectures, seven language classifiers, and five VLMs were evaluated throughout; three generative models performed exploratory note-drafting. Findings: Raw ChatGPT-generated text produced the highest performance of any modality or condition: language models reached near-ceiling AUC (0.992 to 1.000, n=300), exceeding every vision model (AUC 0.799 to 0.880) and every VLM image-only probe (AUC 0.63 to 0.69). On the 297-case pre/post analysis, this advantage depended heavily on label source: language and text-derived VLM performance fell substantially from ChatGPT-only to radiologist-confirmed labels (e.g. BERT-base AUC 0.999 to 0.837), while vision-model performance was stable or modestly improved (e.g. DenseNet-121 0.867 to 0.891). The radiologist reclassified 62 of 297 cases (21%) relative to ChatGPT's provisional read, and a meaningful proportion of raw ChatGPT text was clinically uninterpretable or unsupported by the imaging. Interpretation: As shown here for the first time, raw, image-derived AI-generated text yields the highest apparent classification performance in this benchmark, but this reflects the text's alignment with its own self-generated labels rather than verified diagnostic content, and a substantial share of that text is not clinically explainable. Radiologist-confirmed text and labels give a lower but trustworthy estimate of true performance, on which convolutional neural network (CNN) vision models remain a stable, comparatively inexpensive baseline. Multimodal dental AI should report performance separately by modality and label provenance rather than pooling headline metrics.

5
CT ECV Mapper: an interactive 3D Slicer application with a batch-capable pipeline for voxelwise CT-derived extracellular volume mapping of the liver and hepatic tumors

Suzuki, M.

2026-08-11 radiology and imaging 10.64898/2026.08.09.26360018 medRxiv
Top 0.1%
18.6%
Show abstract

Background. Extracellular volume fraction (ECV) derived from contrast-enhanced CT is a validated marker of hepatic fibrosis and has been reported to differ between hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma. In published work it is obtained from a small number of hand-placed two-dimensional regions of interest, and the software that computes it is either tied to one manufacturer's workstation or based on spectral or dual-energy acquisition. We are not aware of an accessible tool that produces voxelwise liver ECV maps from conventional single-energy multiphase CT. Methods. We developed CT ECV Mapper, a scripted 3D Slicer extension with a three-layer architecture whose numerical core imports neither slicer nor vtk and is unit-tested outside 3D Slicer. The interactive application provides two-stage registration that the operator inspects and accepts before any ECV is computed, operator-placed three-dimensional regions of interest, user-adjustable calculation parameters, a voxelwise ECV color map and ROI statistics; the same logic layer can be driven unattended across a cohort. The tool was applied to the 164 patients of the public WAW-TACE multiphase HCC/TACE dataset that have both unenhanced and delayed-phase series. Results. 156 of 164 cases (95.1%) completed unattended. Whole-liver ECV had a median of 36.2% (interquartile range 31.9-41.5), consistent with published CT-ECV values for fibrotic and cirrhotic liver. Registering the arterial and portal phases on demand extended tumor ECV from the 38 lesions a conventional two-phase pipeline can reach to 248 lesions in 156 patients. Every failure was attributable to an identifiable mechanism: craniocaudal field-of-view mismatch between phases in six cases, aortic calcification within the blood-pool region in one, and in one case a labeling error in the source dataset, in which the series declared as unenhanced proved to be a second reconstruction of the portal venous phase; this was detected by the blood-pool validity check rather than by visual review. Conclusions. Voxelwise CT ECV mapping of the liver and of hepatic tumors is feasible from conventional multiphase CT on an open platform, both interactively and as an unattended batch, with quality-control instrumentation that fails explicitly and diagnosably. This is a technical development and feasibility report; the application has not been evaluated against a reference standard and no claim of clinical validity is made.

6
Cross-Domain Knowledge Transfer from Expert-Annotated Gated CT via Synthetic Ungated CT Improves Coronary Artery Calcium Scoring on CT Attenuation Correction Scans

Shanbhag, A.; Miller, R. J.; Killekar, A.; Marcinkiewicz, A. M.; Zhou, J.; Lemley, M.; Kamagate, A.; Van Kriekinge, S. D.; Kavanagh, P. B.; Feher, A.; Miller, E. J.; Liang, J. X.; Berman, D. S.; Dey, D.; Leahy, R. M.; Slomka, P.

2026-07-06 radiology and imaging 10.64898/2026.07.02.26356002 medRxiv
Top 0.1%
18.3%
Show abstract

Background: Coronary artery calcium (CAC) is an established measure of coronary atherosclerosis from computed tomography (CT). While deep learning (DL) can quantify CAC from non-dedicated CT, the accuracy is limited by image quality. Purpose: We derived and validated a novel method for DL CAC segmentation on ultra-low dose CT attenuation correction (CTAC) scans that is trained with synthetic low-dose, ungated images. Materials and Methods: Models were trained using one center and externally tested in two other centers. Synthetic, ungated CT scans were generated so that expert segmentations from dedicated CAC scans could be used as ground truth for perfectly registered synthetic images through knowledge adaptation (KAD-CAC). We evaluated agreement between CAC scoring methods vs expert readers on a per-patient and per-vessel basis, as well as associations with the primary outcome of death or myocardial infarction (MI). Results: The DL models were externally tested on 5969 patients with a median age of 64 (IQR 56 - 73), of whom 50.2% were male. The KAD-CAC model had higher Cohens kappa K (0.86, 95% CI 0.85 - 0.87) compared to previous convolutional LSTM model (K 0.78, 95% CI 0.76 - 0.80, p<0.01), or models trained with only gated images (K 0.81, 95% CI 0.80 - 0.82, p<0.01). Net reclassification improvement for CAC stratified risk of death or MI, was greatest for the KAD-CAC model over baseline including age, sex, hypertension, diabetes, dyslipidemia, family history, smoking, stress total perfusion deficit, and left ventricular ejection fraction. Conclusion: We use paired synthetic ungated scans to transfer expert gated CAC annotations into the ungated domain, resulting in substantially better vessel-level CAC scoring and improved risk stratification.

7
Chest Radiography AI Concordance and Lung Cancer Linkage in a Large Health Check-up Cohort

Fujita, Y.; Saito, S.; Yagishita, S.; Araya, J.; Nakagawa, R.

2026-07-28 respiratory medicine 10.64898/2026.07.27.26359077 medRxiv
Top 0.1%
18.2%
Show abstract

Purpose To evaluate the implementation characteristics of a commercially available chest radiography artificial intelligence (AI) system in a large real-world health check-up cohort using workflow-level, lesion-specific, and exploratory retrospective lung cancer case analyses. Methods This retrospective single-centre study included 298,991 consecutive health check-up chest radiographs from 114,866 individuals obtained between 2019 and 2023 and interpreted under routine double reading by board-certified radiologists. A commercially available AI system was evaluated using two prespecified thresholds: positivity in any of ten findings for the all-score analysis and positivity for nodule or mass for the nodule-focused analysis, both at a manufacturer-recommended score threshold of 15. Because routine radiologist judgement rather than universal CT or pathologic verification served as the reference framework, the primary analyses were interpreted as radiologist-referenced operational concordance analyses. Results Radiologist-referenced sensitivity and specificity were 72.0% and 79.6%, respectively, in the all-score analysis and 87.1% and 91.8%, respectively, in the nodule-focused analysis. Negative predictive values were 99.0% and 100.0%, respectively. Among 48 histopathologically confirmed lung cancer cases, retrospective timeline analyses showed earlier AI positivity than routine radiologist positivity in a subset of cases. These findings should be interpreted as exploratory observations and do not establish prospective clinical benefit. Conclusion In a large health check-up cohort, chest radiography AI demonstrated stable concordance with routine radiologist judgement and high sensitivity for radiologist-reported pulmonary nodules and masses. Exploratory retrospective analyses showed earlier AI positivity in a subset of histopathologically confirmed lung cancer cases, supporting further prospective evaluation of AI-assisted health check-up workflows.

8
Whole-Lung Ct Radiomics-Based Machine Learning Classification of Nontuberculous Mycobacteria Lung Disease Across Geographically Distinct Cohort

Kanagala, A.; Garcia, B.; Dutt, T. S.; Aguilera, S. M.; Pudhota, A. S.; Panjwani, D. D.; Dukkipati, N.; Gaggar, A.; Naidoo, T.; Jololian, L.; Bhatt, S. P.; Margaroli, C.; Bodduluri, S.

2026-07-09 radiology and imaging 10.64898/2026.06.29.26356713 medRxiv
Top 0.1%
15.7%
Show abstract

BACKGROUND Nontuberculous mycobacterial lung disease (NTM-LD) is highly heterogenous, geographically and etiologically, hindering effective timely identification. Prior CT radiomics studies require manual segmentation of pathology. We developed a whole-lung CT radiomics-based machine learning approach and identified common features across two geographically distinct NTM-LD cohorts. STUDY DESIGN AND METHODS 1,300 chest CT scans from China (871 TB; 429 NTM, Dataset 1) and 173 independent NTM cohort from UAB, US. Whole-lung regions were automatically segmented on each scan, and 85 quantitative radiomic features were extracted using a standardized image-processing pipeline. We evaluated two frameworks to assess model performance and generalizability: (1) training on Dataset 1 with external validation on Dataset 2, and (2) training on the combined cohort. Linear discriminant analysis (LDA) was used as the primary classification method. Cross-cohort concordance analysis was performed to evaluate the reproducibility of radiomic features across datasets. RESULTS In Scenario 1, the LDA classifier trained on Dataset 1 achieved an AUC of 0.79 (95% CI, 0.73-0.84) with high specificity (0.91). On the external UAB cohort, the model achieved an AUC of 0.94 (95% CI, 0.90-0.97). In Scenario 2, the combined cohort model achieved an AUC of 0.81 (95% CI, 0.76-0.85) with improved sensitivity (0.61) and precision (0.82). Feature importance analysis identified 16 features consistently ranked among the top 20 in both scenarios, predominantly texture-based descriptors reflecting distinct parenchymal patterns between mycobacterial species. CONCLUSION Whole-lung CT radiomics enables interpretable NTM-LD classification across geographically distinct populations without manual annotation. Suggesting population-independent parenchymal signatures of NTM-LD.

9
Clinical selectivity and failure modes of automated chest radiograph report evaluation metrics: a cross-dataset analysis of ReXErr-v1 and RadEvalX

Naidu, J.; Muralidharan, S.; Prashani, A.; Baskaradoss, V.

2026-08-12 radiology and imaging 10.64898/2026.08.10.26360043 medRxiv
Top 0.1%
15.6%
Show abstract

Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert error counts. BLEU-4, ROUGE-L and METEOR were assessed in ReXErr-v1; RadEvalX analyses included these plus BERTScore, CheXbert, RadGraph F1 and RadCliQ. Outcomes were ReXErr-v1 pairwise win rate and AUROC for clinical-content versus linguistic errors, and RadEvalX Spearman correlation with clinically significant error count and AUROC for any significant error. Confidence intervals used 10,000 clustered percentile bootstrap resamples; Holm adjustment-controlled multiplicity. Results: ReXErr-v1 paired-sentence win rates were 0.986 for BLEU-4, 0.999 for ROUGE-L and 0.998 for METEOR, but discrimination of clinical-content from linguistic errors was modest (AUROC 0.609-0.620). Penalty magnitude was strongly associated with textual change after adjustment for error type (normalised character edit distance coefficient 0.746; 95% CI 0.705-0.788; P<0.001). In RadEvalX, CheXbert showed the highest correlation with clinically significant errors (rho=0.413; 95% CI 0.223-0.578) and highest AUROC (0.742; 95% CI 0.638-0.836). Conclusions: Near-ceiling sensitivity to textual corruption did not imply sensitivity to clinical significance. CheXbert showed the highest alignment with expert error assessment, although pairwise superiority was not demonstrated over all comparators and performance remained moderate.

10
Learning Diagnostic Proficiency in Robotic-Assisted Bronchoscopy with Integrated Cone-Beam CT: A 680-Lesion Learning Curve Analysis

Aissami, N.; Steinack, C.; Engeli, R.; Baumgartner, P.; Amstitz, A.; Tanner, J.; Clarenbach, C.; Ulrich, S.; Kohler, M.; Gaisl, T.

2026-07-22 respiratory medicine 10.64898/2026.07.21.26358556 medRxiv
Top 0.1%
13.5%
Show abstract

Background. Robotic-assisted bronchoscopy combined with integrated cone-beam computed tomography (RAB+CBCT) enables accurate sampling of peripheral pulmonary lesions (PPLs), but the acquisition of diagnostic proficiency and program-level efficiency remains incompletely characterized. Methods. We conducted a single-center cohort study of consecutive RAB+CBCT (Ion endoluminal system, Cios Spin) procedures performed by two experienced interventional pulmonologists. Strict lesion-level diagnostic yield was the primary outcome. Learning curve cumulative sum (LC-CUSUM) analysis determined operator-specific proficiency, followed by conventional CUSUM monitoring of post-proficiency performance. Secondary outcomes included procedure and intubation times, temporal changes in case complexity, and adverse events. Results. Overall, 427 procedures comprising 680 PPLs were analyzed. Median lesion long-axis diameter was 11 mm, 14.4% had a bronchus sign, and 36.5% of procedures involved multiple lesions. Strict lesion-level diagnostic yield was 85.4% (581/680). LC-CUSUM demonstrated proficiency after 48 and 86 PPLs, respectively; thereafter, both operators maintained acceptable performance without crossing the predefined CUSUM decision limit. Median procedure time decreased from 65 minutes during the first 10 procedures to 38 minutes during the last 10. The median intubation time was 70 minutes and declined significantly with increasing experience. Most indicators of lesion complexity remained stable, while short-axis diameter and bronchus-sign prevalence decreased modestly. Adverse-event frequency declined significantly over time. Conclusion. RAB+CBCT achieved high strict diagnostic yield, with heterogeneous operator-specific learning trajectories within a maturing multidisciplinary program. Diagnostic performance, procedural efficiency, and safety improved despite stable or modestly increasing case complexity. These findings support individualized, outcome-based proficiency assessment and longitudinal monitoring, rather than comparative operator ranking or reliance on fixed procedural-volume thresholds.

11
A 3-Minute Education on the False Positive Paradox Improves Trust Calibration in AI-Assisted Intracranial Aneurysm Detection: A Multinational Randomized Controlled Reader Study

Kim, S. H.; Le Guellec, B.; Rossmueller, P.; Schramm, S.; Boese, L.; Nikoubashman, O.; Kottlors, J.; Lichtenstein, T.; Strotzer, Q.; Meddeb, A.; Ziegelmeyer, S.; Steinhelfer, L.; Prucker, P.; Berberich, C.; Canisius, J.; Kreutzinger, V.; Hartl, F.; Schmitzer, L.; Rosenkranz, E.; Leonhardt, Y.; Beutel, T.-M.; Bitzer, F.; Maegerlein, C.; Boeckh-Behrens, T.; Baum, T.; Makowski, M. R.; Kirschke, J. S.; Bressem, K. K.; Adams, L. C.; Baird, G. L.; Wiestler, B.; Hedderich, D. M.

2026-08-28 radiology and imaging 10.64898/2026.08.25.26361324 medRxiv
Top 0.1%
13.3%
Show abstract

Background Even a highly accurate diagnostic test can yield more false-positive than true-positive findings in low-prevalence settings, which is known as the false positive paradox. Radiologists' unawareness of this paradox may foster automation bias, the tendency to excessively rely on AI outputs. Methods In this prospective, multinational, randomized controlled reader study (DRKS00038740), 34 readers from 10 countries (16 residents, 8 general radiologists or fellows, and 10 neuroradiologists) were randomly assigned to a control group (n = 17) or intervention group (n = 17), stratified by experience level. The intervention group reviewed a short, 3-minute educational video explaining the false positive paradox prior to the reading session. Both groups evaluated 20 TOF-MRA studies with AI-flagged findings (10% true-positive, 90% false-positive). Primary outcomes were acceptance rate of false-positive AI findings and follow-up intensity. These were evaluated using mixed models with crossed random effects for reader and case. Results At baseline, readers vastly overestimated the positive predictive value of AI tools for intracranial aneurysm detection (mean estimate, 62.9%; simulation-based estimate, 15.4% [95% interval, 8.1-28.0%]). The intervention reduced the odds of accepting AI false positives (OR 0.50 [upper 95% confidence bound, 0.95], one-sided p = 0.017), with acceptance probabilities of 12.7% (95% CI, 6.0-25.0%) in the intervention group compared to 22.5% (95% CI, 11.6-39.2%) in the control group. The intervention group exhibited a downward shift in follow-up intensity for false positives (OR 0.47 [upper 95% confidence bound, 0.81]; one-sided p = 0.014), recommending follow-up in 39.2% (120/306) of cases, compared to 54.9% (168/306) in the control group. Conclusion A brief education on the false positive paradox improved trust calibration in AI-assisted intracranial aneurysm detection. Our findings highlight the potential of reader-side cognitive debiasing strategies to improve trust calibration and support safer use of AI in radiology.

12
A Real-World Evaluation of Failure Detection for Liver CT Segmentation

Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.

2026-06-29 radiology and imaging 10.64898/2026.06.26.26356692 medRxiv
Top 0.1%
13.0%
Show abstract

Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.

13
Deep Learning based Quantification of Root Exposure in Lower Anterior Teeth using Intraoral camera image

Choi, H.; Choi, Y.-H.; Park, E. Y.; Kang, S.; Kim, E.-K.

2026-07-04 dentistry and oral medicine 10.64898/2026.07.02.26357104 medRxiv
Top 0.1%
12.2%
Show abstract

This study compared and evaluated two widely used deep learning-based artificial intelligence (AI) models, U-Net++ and YOLOv11, for quantifying tooth and root exposure on intraoral camera images of the mandibular anterior lingual region. Intraoral images of the mandibular anterior lingual region were collected from 291 patients (mean age, 52.8 years) at a university hospital dental clinic with institutional review board approval (YUMC IRB 2021-07-019-002). A total of 266 eligible images (mean, 5.50 teeth per image; 3.70 teeth with root exposure) were annotated. YOLOv11 and U-Net++ were fine-tuned using five-fold cross-validation with data augmentation. Model performance was evaluated on a held-out test set of 40 images using Dice coefficient, Intersection over Union (IoU), accuracy, mean Average Precision at an IoU threshold of 0.5 (mAP50), Lins concordance correlation coefficient (CCC), and intraclass correlation coefficient (ICC). Confidence intervals were estimated using 10,000 bootstrap iterations. For tooth segmentation, U-Net++ demonstrated superior performance, with high accuracy (0.981), Dice coefficient (0.971), and IoU (0.944). In contrast, for root segmentation, YOLOv11 outperformed U-Net++, achieving higher Dice (0.860 vs. 0.746) and IoU (0.762 vs. 0.631). Notably, YOLOv11 showed stronger agreement with the ground truth for quantifying the exposed root ratio (ERR) (CCC, 0.973; ICC, 0.975). These findings suggest that accurate detection of root exposure is important for assessing periodontal tissue loss and that YOLOv11 is a promising model for root exposure quantification in intraoral images. YOLOv11-based quantification of root exposure may serve as a useful adjunctive AI tool for screening and monitoring periodontal conditions and may support individualized treatment planning.

14
Beyond BMI: an interpretable integrated body composition index from low-dose chest CT for all-cause mortality risk stratification: a multicentre study

Yi, J.; Patel, K. K.; Miller, R. J. H.; Marcinkiewicz, A. M.; Kamagate, A.; Shanbhag, A.; Hijazi, W.; Lemley, M.; Zhou, J.; Liang, J. X.; Ramirez, G.; Mostafavi, S.; Urs, M.; Spielvogel, C. P.; Slipczuk, L.; Travin, M.; Alexanderson, E.; Caraval-Juarez, I.; Packard, R. R.; Al-Mallah, M.; Ruddy, T. D.; Einstein, A. J.; Feher, A.; Miller, E. J.; Acampa, W.; Knight, S.; Le, V. T.; Mason, S.; Calsavara, V. F.; Chareonthaitawee, P.; Wopperer, S.; Kwan, A. C.; Wang, L.; Li, D.; Fishman, E. K.; Lopez-Ramirez, F.; Berman, D. S.; Kwiecinski, J.; Dey, D.; Di Carli, M. F.; Slomka, P.

2026-08-10 radiology and imaging 10.64898/2026.08.05.26359437 medRxiv
Top 0.1%
12.1%
Show abstract

Background: Body composition is recognized as a major determinant of health outcomes, but its multidimensional nature makes clinical adoption challenging. We sought to develop and validate a body composition index (BCI) for all-cause mortality risk assessment, integrating variables of six body composition tissues. Methods: We analyzed 28509 consecutive patients undergoing myocardial perfusion imaging with routine low-dose chest CT attenuation correction (CTAC) scans acquired during myocardial perfusion imaging (MPI) at 12 centers across four countries. An artificial intelligence-based BCI was developed in a cohort of 15037 patients CTACs by integrating the CT-derived metrics of bone, skeletal muscle, and four adipose tissue compartments, coronary artery calcium score, and basic demographic variables (age, sex, BMI). The performance of BCI for mortality prediction was validated in an internal cohort of 6444 patients and an external cohort of 7028 patients by prognosis, calibration, net benefit, and explainability. Model-based simulation of tissue metrics modification was performed to evaluate estimated mortality risk reduction. Findings: During a median of 3.5 (IQR [1.9, 5.1]) years, 4697 (16%) patients died. In the external testing cohort, the BCI demonstrated excellent discrimination for mortality (area under receiver operating characteristic curve 0.78 (95% CI [0.76, 0.79]) and Harrell concordance index 0.75 [0.73, 0.76]), calibration, and net benefit overall and across pre-specified subgroups stratified by patient characteristics and imaging protocols. Visceral adipose tissue attenuation was the most influential body composition measure, followed by skeletal muscle volume. Simulated improvement in body composition was associated with significant mortality risk reduction. Interpretation: An index combining six body composition measures obtained opportunistically from routine chest CT provides robust mortality risk stratification. By converting complex body composition information into a single interpretable score, the BCI can facilitate clinical implementation of opportunistic CT biomarkers and guide individualized preventive strategies.

15
Artificial Intelligence-Enabled Detection of Vascular Perfusion Defects on Ventilation/Perfusion (V/Q) Scintigraphy for Pulmonary Embolism

Jabbarpour, A.; Moulton, E.; Kaviani, S.; Zeng, W.; Ghassel, S.; Akbarian, R.; Couture, A.; Roy, A.; Liu, R.; Al-ali, Y.; Foufa, Y.; Hejji, N.; AlSulaiman, S.; Shirazi, Z.; Leung, E.; Klein, R.

2026-07-08 radiology and imaging 10.64898/2026.06.25.26356599 medRxiv
Top 0.1%
11.9%
Show abstract

Accurate interpretation of planar ventilation-perfusion (V/Q) scintigraphy, used for diagnosing pulmonary embolism (PE) based on PIOPED/EANM guidelines, requires objective assessment of mismatched V/Q defects. Manual delineation of V/Q defects is time-consuming, subject to interobserver variability, and rarely performed in practice, limiting standardized reporting and quantification of disease burden. To address these challenges, we evaluated four modern AI models for automated segmentation of vascular perfusion defects in planar V/Q scans and compared their performance to human annotators. We retrospectively identified 2,118 patients who underwent planar V/Q scans at The Ottawa Hospital (June 2019-February 2023). Six standard projections (ANT, POST, LAO, RAO, LPO, RPO) were included. Four 2D neural networks (U-Net, nnU-Net, Swin UNETR, and a Bottleneck Transformer U-Net [BTU-Net]) were trained on 1,313 patients (7,878 projections) and validated on 329 (1,974 projections) using physician-annotated defects. A hold-out test set of 46 high probability patients was used to evaluate segmentation quality, and defect detection accuracy using free-response receiver operating characteristic (FROC) analysis, where BTU-Net was the only model performing on par with human readers, showing robust sensitivity across the entire range of segmentation probabilities. At 1.5 false positives per projection rate (FPPR), BTU-Net outperformed other models with a sensitivity of 0.529 {+/-} 0.026, On a separate hold-out set of low likelihood of disease patients (n=430), the lowest FPPR was 0.08 {+/-} 0.01 for BTU-Net (P<0.0001). BTU-Net enables rapid, consistent, and accurate interpretation of planar V/Q scans. Such tools may enhance diagnostic efficiency, standardize reporting, and support non-expert readers in evaluating PE.

16
Human In the Loop Challenges for Quality Annotation of Pre-Cancer Lesions in Clinical Oral Images

Mandal, S.; Mendonca, P.; Gurushanth, K.; Thakur, H.; Birur, P.; Shetty, A.; Pal, D.

2026-07-04 dentistry and oral medicine 10.64898/2026.07.02.26355859 medRxiv
Top 0.1%
11.9%
Show abstract

Background: The hyperplasia and dysplasia stage (pre-cancer) offers a viable opportunity to reduce the incidence and mortality of oral cancer through early prevention. Smartphone-based Artificial Intelligence (AI) enabled screening of potentially malignant oral lesions offers a scalable solution for this in resource-constrained settings. However, developing accurate and explainable AI segmentation models require high-quality, pixel-level annotated data. This process that is prohibitively expensive, time-consuming, and prone to inter-observer subjectivity among clinical experts. Methods: We designed and empirically validated a Deep Learning-driven Human-in-the-Loop (HITL) framework to pixel-annotate a dataset of 3026 clinical oral images. Using an iterative pseudo-labeling pipeline, we evaluated the model's learning dynamics and performance evolution across five training cycles. We conducted controlled experiments to quantify the networks tolerance to intermediate level of label noise (unreviewed pseudo-labels) to resolve clinical subjectivity using pixel-wise Cohen's Kappa and the STAPLE consensus algorithm. Results: Iterative self-training produced sustained improvements in lesion detection and spatial localization. However, controlled experiments revealed that including even a modest fraction ({approx}10%) of unreviewed pseudo-labels led to a three-to-four-fold increase in training convergence instability and induced a conservative prediction bias that negatively impacted model recall. When measured against multi-expert ground truth, the model's performance converged with the inter-rater reliability ceiling ({kappa} {approx} 0.65), indicating that its predictions fell within the envelope of human agreement. Conclusions: Our findings emphasize that a final expert-driven quality assurance step remains absolutely essential to mitigate training instability, confirmation bias, and clinically unacceptable drops in recall caused by label noise. Overall, this work provides a scalable, empirically validated blueprint for building domain-specific medical imaging datasets in low-resource global health settings, where the dual challenges of annotation cost and inter-observer variability are most acute.

17
Algorithmic implementation of pancreatic cancer staging guidelines: comparison with a retrieval-augmented large language model

Komaba, A.; Amakawa, A.; Tozuka, R.; Sato, J.; Fujihara, K.; Emoto, M.; Sawada, S.; Kasai, S.; Sakamoto, K.; Shimura, K.; Johno, Y.; Nakamoto, K.; Ichikawa, S.; Johno, H.

2026-07-02 radiology and imaging 10.64898/2026.06.30.26356912 medRxiv
Top 0.1%
10.0%
Show abstract

Purpose: To implement a comprehensive knowledge-based algorithm (KBA) for pancreatic cancer staging based on the current Japanese guidelines and to evaluate its performance as a clinical decision support system in comparison with a retrieval-augmented large language model (LLM) system. Materials and methods: A KBA covering TNM classification, stage classification, and resectability classification was implemented as a web application. The correctness of the system outputs was exhaustively verified for all possible inputs. Subsequently, six non-board-certified radiologists performed pancreatic cancer staging for 12 simulated cases with imaging findings under three conditions: unassisted, LLM-assisted, and KBA-assisted. Staging accuracy and staging time were compared among the three conditions using pairwise proportion z-tests and Welch's t-tests, respectively. Results: In the comparative experiment, staging accuracy was 81.9%, 80.6%, and 98.6% in the unassisted, LLM-assisted, and KBA-assisted conditions, respectively. Mean staging time was 229.2, 401.9, and 196.2 s, respectively. The KBA-assisted condition showed higher accuracy than both the unassisted and LLM-assisted conditions (both p<0.001). Staging time was longer in the LLM-assisted condition than in the other two conditions (both p<0.001). Conclusion: A comprehensive KBA for pancreatic cancer staging based on the current Japanese guidelines was implemented and exhaustively verified. In a preliminary comparative experiment, KBA assistance improved staging accuracy without increasing staging time, whereas LLM assistance increased staging time without improving staging accuracy. These findings suggest that verified KBA systems may be feasible and useful for clinical tasks governed by explicit guideline-based rules.

18
Benchmarking Open-Source Vision-Language Models for Brain Metastasis Assessment on Single-Slice Contrast-Enhanced MRI

Kim, J.; Kim, B.-s.; Ko, J. S.; Dong, J.; Youn, S. Y.; Jang, J.; Ahn, K.-J.

2026-08-26 radiology and imaging 10.64898/2026.08.24.26361169 medRxiv
Top 0.1%
9.7%
Show abstract

Purpose Open-source vision-language models (VLMs) can be locally deployed without external internet access, potentially enhancing data security. This study compared the diagnostic performance of general-purpose and medical-purpose open-source VLMs and evaluated their ability to characterize brain metastases on contrast-enhanced (CE) MRI. Materials and Methods Sixty lesion-positive axial CE T1-weighted images and sixty matched lesion-negative images from 60 patients were analyzed using three general-purpose VLMs-InternVL3-8B, Qwen2.5-VL-7B-Instruct, and MiniCPM-V-4.5-and three medical-purpose VLMs-MedGemma-4B-it, LLaVA-Med v1.5, and HuatuoGPT-Vision-7B. Lesion detection performance was assessed using sensitivity, specificity, and balanced accuracy. On lesion-positive images, accuracy was evaluated for lesion count, laterality, anatomic location, enhancement pattern, necrosis, vasogenic edema, and mass effect. Model differences were assessed using Cochran's Q tests followed by pairwise McNemar tests with Benjamini-Hochberg correction. Results The median age of the study patients was 67 years (IQR, 61.0-70.5 years), and 35 patients were male (58.3%). MiniCPM-V-4.5 showed the most balanced diagnostic performance, with a sensitivity of 78.3% (95% CI, 66.4-86.9%) and a specificity of 85.0% (95% CI, 73.9-91.9%), and significantly higher balanced accuracy than all other models. Significant overall differences were observed for lesion count, laterality, location, enhancement pattern, necrosis, and mass effect, but not for vasogenic edema (FDR-adjusted P = 0.056). HuatuoGPT-Vision-7B and MedGemma-4B-it showed relatively consistent accuracy across multiple image assessment tasks, although their performance remained modest. Conclusion Our study demonstrated substantial heterogeneity in the performance of open-source VLMs in brain metastasis evaluation, and medical-purpose VLMs did not outperform general-purpose VLMs.

19
Acute Ischemic Stroke Detection on Non-Contrast CT: A Deep Learning Approach

Goyal, A.; Stevens, R. D.

2026-06-23 radiology and imaging 10.64898/2026.06.20.26356152 medRxiv
Top 0.1%
8.7%
Show abstract

Acute ischemic stroke (AIS) is a leading cause of disability and death while effective treatment requires quick and accurate diagnosis. Non-contrast CT (NCCT) is widely used in the initial screening of AIS, but stroke detection is challenging because early changes on NCCT are subtle or indistinguishable. Using hyperacute NCCTs as inputs and diffusion-weighted MRI as ground truth, we trained a deep learning algorithm to classify patients with AIS and segment the stroke lesions. We hypothesized that this approach would accurately detect hyperacute tissue density changes on NCCT. For the classification task, our ResNet50 model delivered the best performance (with 98.5% accuracy, 97.4% precision, and 100% recall on an evaluation set). Classification performance remained strong when restricted to lesions smaller than 5 mL, which constituted the majority of our evaluation cases. For the segmentation task accomplished using a range of U-Net architectures, performance was acceptable for large lesions and declined sharply for smaller lesions. Together, these findings demonstrate the feasibility of deep learning for AIS detection and represent a step towards faster triage and treatment for stroke patients.

20
Automated Airways Characterization and Assessment of Cystic Fibrosis from CT Imaging

Chong Chie, J. A. K. H.; Cooper, M. L.; Persohn, S. A.; Burton, C. P.; Salama, P.; Territo, P. R.

2026-06-18 radiology and imaging 10.64898/2026.06.09.26355170 medRxiv
Top 0.1%
8.3%
Show abstract

Background Advancements in medical imaging have enabled non-invasive diagnosis and staging of cystic fibrosis (CF) using CT scans, revealing dilated airways, an increased number of visible airways, and airway generation splits in these patients. However, manual characterization of airways remains time-consuming and challenging due to the numerous structural changes, thereby limiting clinical feasibility. This study aims to develop an automated algorithm to characterize airways from segmented lung CT scans and apply this to a retrospective population. This approach reduces the time required to analyze images and obtain disease-staging results. Methods This framework consists of two stages. The first stage extracts and skeletonizes the airway tree from lung CTs, while the second stage measures lung features, including airway volumes, branch counts, generation splits, diameters, and cross-sectional areas. This permits comprehensive characterization for use in clinical assessment. Results The airways analysis was performed on 169 CT volumes ranging in age from 6 to 18 years of age, revealing substantial differences in detected airway branches, generation splits, and normalized airway volume between the control and CF groups. The framework also measures airway diameters and cross-sectional areas, revealing an increase in the number of small airways in cystic fibrosis patients, due to early bronchiectasis. These findings align with previous research and demonstrate the framework's ability to accurately quantify airway changes in patients with CF. Discussion The framework extracts entire airway trees, facilitating measurements of volume, branch count, diameters, and cross-sectional areas, which change with CF severity and/or treatment. However, partial lung atelectasis can limit the accuracy of airway detection in moderate-to-severe cases. Funding NIA U54 AG054345 and NIA R21 AG07857501